Papers with vision-language model
ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT (2025.acl-short)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) has improved translation by using Transformer-based models, but still struggles with word ambiguity and context. |
| Approach: | They create a new Czech-to-polish e-commerce product translation dataset coupled with images and product metadata consisting of 11,400 sentence pairs. |
| Outcome: | The proposed model incorporates visual cues alongside textual data to improve translation quality. |
Simplifying Outcomes of Language Model Component Analyses with ELIA (2026.eacl-demo)
Copied to clipboard
| Challenge: | ELIA is an interactive web application that simplifies the outputs of various language model component analyses for a broader audience. |
| Approach: | They propose to use a vision-language model to automatically generate natural language explanations for the complex visualizations produced by these methods. |
| Outcome: | The proposed system integrates three key techniques and generates natural language explanations for complex visualizations. |
Detecting AI-Generated Content on Social Media with Multi-modal Language Models (2026.acl-industry)
Copied to clipboard
Chenyang Yang, Shen Yan, Yibo Yang, Litao Hu, Yuchen Liu, Yuan Zeng, Hanchao Yu, Yinan Zhu, Sumedha Singla, Brian Vanover, Huijun Qian, Zihao Wang, Fujun Liu, Aashu Singh, Jianyu Wang, Xuewen Zhang
| Challenge: | Existing methods for AI-generated content detection face poor generalization to newer models, reliance on single modalities, and lack of interpretable explanations. |
| Approach: | They propose a model that curates diverse social media data and trains a vision-language model for detection and explanation. |
| Outcome: | The proposed model achieves state-of-the-art detection performance on public benchmarks and observes positive downstream impacts on user engagement. |
Optical Character Recognition for the International Phonetic Alphabet (2026.eacl-short)
Copied to clipboard
| Challenge: | Grammar books are increasingly used as additional reference resources for low-resource languages . a significant portion of these documents come from scans and require an OCR tool . |
| Approach: | They compare two neural OCR frameworks and a large vision-language model with a synthetic dataset based on Wiktionary to study the International Phonetic Alphabet (IPA). |
| Outcome: | The proposed model improves on the International Phonetic Alphabet (IPA) character set. |
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent advances in vision-language models have accelerated research into models capable of advanced reasoning based on images. |
| Approach: | They propose a method that leverages vision-language models to convert charts into table format . they use Large Language Model (LLM) for reasoning to extract only the essential information . |
| Outcome: | The proposed method extracts only the elements necessary for chart reasoning without the need for additional annotations or datasets. |
VLStereoSet: A Study of Stereotypical Bias in Pre-trained Vision-Language Models (2022.aacl-main)
Copied to clipboard
| Challenge: | Existing studies on pre-trained vision-language models have focused on measuring biases and stereotypes in a single modality. |
| Approach: | They extend a recently released stereotypical bias dataset into a vision-language probing dataset called VLStereoSet to measure stereotypical biased vision-linguistic models. |
| Outcome: | The proposed probing task measures stereotypical bias in vision-language models and its intra-modal and inter-modal biases. |
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing methods that optimize for relevance overlook document trustworthiness . Generative information retrieval (GenIR) is a promising paradigm for retrieval tasks . |
| Approach: | They propose an Authority-aware Generative Retriever (AuthGR) that incorporates authority into GenIR. |
| Outcome: | The proposed framework improves authority and accuracy in real-world user engagement and reliability. |
One More Modality: Does Abstract Meaning Representation Benefit Visual Question Answering? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | incorporating explicit semantic information, in the form of Abstract Meaning Representation graphs, can enhance VQA models. |
| Approach: | They augment two vision-language models with sentence- and document-level AMRs . they find that in well-resourced settings, models are negatively impacted by AMR . |
| Outcome: | The proposed model improves in well-resourced and low-resource settings with AMR graphs . the model achieves 13.1% relative gain using sentence-level AMRs compared with the smaller model . |
DE-CLIP: Few-Shot Anomaly Detection via Difference-Guided Embedding Editing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to detect anomalies are limited due to the lack of anomalous samples . |
| Approach: | They propose a framework that edits text embeddings based on the differences between normal and anomalous samples. |
| Outcome: | The proposed framework achieves 96.6% and 96.99% AUROC on MVTec datasets. |
Graph-guided Cross-composition Feature Disentanglement for Compositional Zero-shot Learning (2025.findings-acl)
Copied to clipboard
Yuxia Geng, Runkai Zhu, Jiaoyan Chen, Jintai Chen, Xiang Chen, Zhuo Chen, Shuofei Qiao, Yuxiang Wang, Xiaoliang Xu, Sheng-Jun Huang
| Challenge: | Disentanglement of visual features of primitives (i.e., attributes and objects) has shown exceptional results in Compositional Zero-shot Learning (CZSL). |
| Approach: | They propose a solution that takes multiple compositions as inputs and constrains disentangled primitive features to be general across compositions. |
| Outcome: | The proposed architecture significantly improves performance on three popular CZSL benchmarks and has been verified by solid ablation studies. |
Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video Understanding (2025.acl-long)
Copied to clipboard
| Challenge: | Existing multimodal large language models (LLMs) have shown impressive performance on the video understanding task, but extremely long videos still pose significant challenges to their context length, memory consumption, and computational complexity. |
| Approach: | They propose a vision-language model named Sophia for long video understanding which can efficiently handle hour-scale long videos. |
| Outcome: | The proposed model exhibits competitive performance compared to existing video understanding baselines across various benchmarks for long video understanding with reduced time and memory consumption. |
VCD: A Dataset for Visual Commonsense Discovery in Images (2025.findings-acl)
Copied to clipboard
| Challenge: | Visual commonsense data sets lack visual grounded representations of commonsensense . existing knowledge bases lack visual-based knowledge tied to actual visual scenes . |
| Approach: | They present a large-scale visual commonsense dataset with over 100,000 images and 14 million object-commonsense pairs that integrates both Seen (directly observable) and Unseen (inferrable) commonsens. |
| Outcome: | The proposed model integrates Seen (directly observable) and Unseen (inferrable) commonsense across Property, Action, and Space aspects. |
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Among the minority groups under-represented in AI, data from low-income households are often overlooked in data collection and model evaluation. |
| Approach: | They evaluate the performance of a vision-language model on a geo-diverse dataset . they highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development. |
| Outcome: | The proposed model performs lower for the poorer groups than the wealthier groups across topics and countries. |
ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a vision-language model with commonsense knowledge can reason beyond common sense . however, pre-trained vision-linguistic models are incapable of interpreting counter-intuitive content . |
| Approach: | They introduce a probing dataset to evaluate vision-language models' reasoning abilities . they use images that defy commonsense knowledge to test their reasoning abilities. |
| Outcome: | The proposed dataset evaluates whether pre-trained vision-language models can reason beyond common sense . it contains images that defy commonsense knowledge with regards to color, shape, material, size and position . |
Cross-Modal Taxonomic Generalization in (Vision-) Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that language models learn from surface form to learn from more grounded evidence. |
| Approach: | They propose to use a vision-language model to learn hypernyms from images . they find that the model can recover this knowledge and generalize even when there is no hypernomia in the image. |
| Outcome: | The proposed model can recover this knowledge and generalize even when the model receives no evidence of hypernyms during training. |
Aligning Large Multimodal Models with Factually Augmented RLHF (2024.findings-acl)
Copied to clipboard
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liangyan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, Trevor Darrell
| Challenge: | Large Multimodal Models (LMMs) are built across modalities and the misalignment between two modality can result in "hallucination" . developing LMMs faces challenges such as a lack of data and a limited number of data sets. |
| Approach: | They propose a new algorithm that augments the reward model with additional factual information such as image captions and ground-truth multi-choice options. |
| Outcome: | The proposed approach improves on the LLaVA-Bench dataset with the 96% performance level of the text-only GPT-4 and an improvement of 60% on MMHAL-BENCH over other baselines. |
EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained vision-language models have achieved impressive results in a range of vision-linguistic tasks. |
| Approach: | They propose a distilling then pruning framework to compress large vision-language models into smaller, faster ones. |
| Outcome: | The proposed framework reduces the size of a pre-trained large vision-language model and improves its performance on vision-linguistic tasks. |
FigEx: Aligned Extraction of Scientific Figures and Captions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | FigEx is a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Approach: | They propose a vision-language model to extract aligned pairs of subfigures and subcaptions from scientific papers. |
| Outcome: | The proposed model improves subfigure detection APb over Grounding DINO by 0.023 and boosts caption separation BLEU over Llama-2-13B by 0.465. |
TransferCVLM: Transferring Cross-Modal Knowledge for Vision-Language Modeling (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent large vision-language multimodal models pre-trained with huge amount of image-text pairs show remarkable performances in downstream tasks. |
| Approach: | They propose a method of efficient knowledge transfer that integrates pre-trained uni-modal models into a combined vision-language model without pre-training . they propose to fine-tune the model and transfer multimodal knowledge from a teacher vision-linguistic model to the CVLM for each task application. |
| Outcome: | The proposed method outperforms existing vision-language models in downstream tasks. |
Think before Go: Hierarchical Reasoning for Image-goal Navigation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for image-goal navigation fail to extract informative visual cues, leading agents to wander around. |
| Approach: | They propose a framework that decomposes image-goal navigation into high-level planning and low-level execution. |
| Outcome: | The proposed method is superior to existing methods in both simulation and real-world environments. |
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on explainable evaluation metrics generate explanations without standardized criteria and the overall quality of the generated explanations remains unverified. |
| Approach: | They propose a reference-free evaluation metric that provides structured explanations based on fluency, relevance, and descriptiveness. |
| Outcome: | The proposed evaluation template achieves state-of-the-art on benchmark datasets while providing significantly higher-quality explanations than existing metrics. |
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling (2025.acl-long)
Copied to clipboard
| Challenge: | Chart generation requires strong visual design skills and precise coding capabilities that embed the desired visual properties into code. |
| Approach: | They propose a vision-language model-based multi-agent framework for effective automatic chart generation. |
| Outcome: | The proposed framework achieves a 5.2% improvement in the F1 score over the current best chart generation task. |
Glyph: Scaling Context Windows via Visual-Text Compression (2026.acl-long)
Copied to clipboard
Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang, Hongning Wang, Minlie Huang
| Challenge: | Large language models (LLMs) traditionally represent text as sequences of discrete tokens . a long-context scaling problem requires processing more tokens more efficiently . |
| Approach: | They propose a framework that renders long texts into compact visual pages and processes them with a vision-language model. |
| Outcome: | The proposed framework renders long texts into compact visual pages and processes them with a vision-language model. |